Papers with visual generation
Colorism in Multimodal AI: An Empirical Exploration of Socioeconomic Linguistic Bias in Text-to-Image Generation (2026.eacl-srw)
Copied to clipboard
| Challenge: | Socioeconomic inequalities worldwide are deeply linked to ethnoracial hierarchies and stereotypes, argues a new study. |
| Approach: | They use a Monk Skin Tone scale to benchmark VLMs and annotators . they then use linguistic cues to vary skin-tone representations in text-to-image generation . |
| Outcome: | The study compares 3 small VLMs and 60 human annotators on the monk skin tone scale with 210 occupations and produces over 2,500 portraits across 3 large VLM models. |
UniFashion: A Unified Vision-Language Model for Multimodal Fashion Retrieval and Generation (2024.emnlp-main)
Copied to clipboard
| Challenge: | e-commerce tasks such as multimodal retrieval and multimodal generation are largely ignored due to the diversity of the multimodal fashion domain. |
| Approach: | They propose a framework that integrates image generation with retrieval and text generation tasks. |
| Outcome: | The proposed framework outperforms state-of-the-art models across fashion tasks. |